|
|
|
| Visual-Tactile Multimodal Learning via Hierarchical Tactile-Guided Spatial Attention and Dynamic Temporal Fusion |
| FAN Chenglong1, HU Lihua1, HU Jianhua2 |
1. School of Computer Science and Technology, Taiyuan University of Science and Technology, Taiyuan 030024; 2. Engineering Laboratory for Intelligent Industrial Vision, Institute of Automation, Chinese Academy of Sciences, Beijing 100190 |
|
|
|
|
Abstract Visual and tactile modalities differ in perceptual scope and information form. Existing methods fail to effectively associate tactile contact information with local visual regions. Moreover, fixed fusion strategies cannot adapt to the dynamic variations in the contributions of the two modalities across different interaction stages. To address these issues, a visual-tactile multimodal learning method via hierarchical tactile-guided spatial attention and dynamic temporal fusion(HTA-DTF) is proposed. First, tactile features are utilized as queries to guide the visual branch to focus on contact-related regions, while intermediate- and high-level visual features are integrated to capture both local texture details and high-level semantic information. Second, a dual-stream Mamba architecture is employed to model the temporal dependencies of visual and tactile sequences, and a dynamic gating mechanism is adopted to adaptively adjust the fusion proportions of the two modalities across different interaction stages. Finally, temporal attention pooling performs weighted aggregation over the fused sequence, emphasizing key interaction moments while suppressing interference from redundant time steps. Experiments on the Touch and Go material recognition dataset demonstrate that HTA-DTF achieves high material recognition accuracy. Ablation studies and visualization analyses further verify the effectiveness of the proposed components.
|
|
Received: 20 May 2026
|
|
|
| Fund:National Natural Science Foundation of China(No.62273248,62573406) |
|
Corresponding Authors:
HU Lihua, Ph.D., professor. Her research interests include computer vision, artificial intelligence, and pa-ttern recognition.
|
About author:: FAN Chenglong, Master student. His research interests include robotic intelligent per-ception, visual-tactile fusion, and multimodal deep learning. HU Jianhua, Ph.D., professor. His research interests include intelligent robotics and 3D vision. |
|
|
|
[1] Lederman S J, Klatzky R L.Haptic perception: a tutorial[J]. Atten-tion, Perception, & Psychophysics, 2009, 71(7): 1439-1459. [2] Luo S, Bimbo J, Dahiya R, et al. Robotic tactile perception of object properties: a review[J]. Mechatronics, 2017, 48: 54-67. [3] Calandra R, Owens A, Upadhyaya M, et al. The feeling of success: does touch sensing help predict grasp outcomes[C]//Proceedings of the 1st International Conference on Robot Learning. San Diego, USA: JMLR, 2017: 314-323. [4] Lee M A, Zhu Y K, Srinivasan K, et al. Making sense of vision and touch: self-supervised learning of multimodal representations for contact-rich tasks[C]//Proceedings of the IEEE International Conference on Robotics and Automation. Washington, USA: IEEE, 2019: 8943-8950. [5] Cui S W, Wang R, Wei J H, et al. Self-attention based visual-tac-tile fusion learning for predicting grasp outcomes[J]. IEEE Robotics and Automation Letters, 2020, 5(4): 5827-5834. [6] Yang F Y, Ma C Y, Zhang J Z, et al. Touch and Go: learning from human-collected vision and touch[C]//Proceedings of the 36th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2022: 8081-8103. [7] Kerr J, Huang H, Wilcox A, et al. Self-supervised visuo-tactile pre-training to locate and follow garment features[EB/OL].[2026-04-17]. https://arxiv.org/pdf/2209.13042. [8] Dave V, Lygerakis F, Rueckert E.Multimodal visual-tactile representation learning through self-supervised contrastive pre-training[C]//Proceedings of the IEEE International Conference on Robotics and Automation. Washington, USA: IEEE, 2024: 8013-8020. [9] Wu Z Y, Zhao Y Q, Luo S.ConViTac: aligning visual-tactile fusion with contrastive representations[C]//Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. Wa-shington, USA: IEEE, 2025: 8545-8552. [10] Lin T Y, Dollar P, Girshick R, et al. Feature pyramid networks for object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2017: 936-944. [11] Xu Z Q J, Zhang Y Y, Luo T, et al. Frequency principle: Fourier analysis sheds light on deep neural networks[J]. Communications in Computational Physics, 2020, 28(5): 1747-1767. [12] Anderson P, He X D, Buehler C, et al. Bottom-up and top-down attention for image captioning and visual question answering[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2018: 6077-6086. [13] Hochreiter S, Schmidhuber J.Long short-term memory[J]. Neural Computation, 1997, 9(8): 1735-1780. [14] Bai S J, Kolter J Z, Koltun V.An empirical evaluation of generic convolutional and recurrent networks for sequence modeling[EB/OL]. [2026-04-17].https://arxiv.org/pdf/1803.01271. [15] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2017: 6000-6010. [16] Doherty J, Gardiner B, Siddique N, ,et al. A novel visuo-tactile object recognition pipeline using transformers with feature level fusion[C/OL]//Proceedings of the International Joint Conference on Neural Networks. Washington. A novel visuo-tactile object recognition pipeline using transformers with feature level fusion[C/OL]//Proceedings of the International Joint Conference on Neural Networks. Washington, USA: IEEE, 2024. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10650147. [17] Jiang C P, Xu W Q, Li Y T, ,et al. Capturing forceful interaction with deformable objects using a deep learning-powered stretchable tactile array[J/OL]. Nature Communications, 2024, 15. https://www.nature.com/articles/s41467-024-53654-y.pdf. [18] Gu A, Goel K, Re C.Efficiently modeling long sequences with structured state spaces[EB/OL]. [2026-04-17].https://arxiv.org/pdf/2111.00396v3. [19] Gu A, Dao T.Mamba: linear-time sequence modeling with selective state spaces[EB/OL]. [2026-04-17].https://arxiv.org/pdf/2312.00752. [20] Liu Y, Tian Y J, Zhao Y Z, et al. VMamba: visual state space model[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2024: 103031-103063. [21] Hatamizadeh A, Kautz J.MambaVision: a hybrid Mamba-Transformer vision backbone[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2025: 25261-25270. [22] Arevalo J, Solorio T, Montes-Y-Gómez M, et al. Gated multimodal networks[J]. Neural Computing and Applications, 2020, 32(14): 10209-10228. [23] Cao G Q, Zhou Y, Bollegala D, et al. Spatio-temporal attention model for tactile texture recognition[C]//Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. Washington, USA: IEEE, 2020: 9896-9902. [24] Ilse M, Tomczak J, Welling M.Attention-based deep multiple instance learning[C]//Proceedings of the 35th International Confe-rence on Machine Learning. San Diego, USA: JMLR, 2018: 2127-2136. [25] GAO R H, DOU Y M, LI H, et al. The ObjectFolder benchmark: multisensory learning with neural and real objects[C]//Procee-dings of the IEEE/CVF Conference on Computer Vision and Pa-ttern Recognition. Washington, USA: IEEE, 2023: 17276-17286. [26] Deng J, Dong W, Socher R, et al. ImageNet: a large-scale hierarchical image database[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2009: 248-255. [27] Loshchilov I, Hutter F.Decoupled weight decay regularization[EB/OL]. [2026-04-17].https://arxiv.org/pdf/1711.05101. [28] Srivastava N, Hinton G, Krizhevsky A, et al. Dropout: a simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 2014, 15: 1929-1958. [29] Loshchilov I, Hutter F.SGDR: stochastic gradient descent with warm restarts[EB/OL]. [2026-04-17].https://arxiv.org/pdf/1608.03983. |
|
|
|